September 1, 2026 Anthropic Trains a Model to Be Bad on Purpose — and Watches What Happens Research Anthropic Alignment Reward Hacking